General Language Model Evaluation on OlmoBaseEval HeldOut
33.7LBPP ScoreNemo. 3 Nano
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Nemo. 3 Nano# Parameters=30B A3B, # Train Tokens=25T, Train Compute (10^23 FLOPs)=5.42026.04 | 33.7 | 78.2 | 53.5 | 32.8 | |
| Kimi Linear# Parameters=48B A3B, # Train Tokens=6T, Train Compute (10^23 FLOPs)=1.02026.04 | 31.1 | 68.6 | 50.7 | 40.8 | |
| Falcon H1# Parameters=7B, # Train Tokens=12T, Train Compute (10^23 FLOPs)=5.52026.04 | 30.2 | 75.5 | 50 | 34.3 | |
| Nemotron-H# Parameters=8B, # Train Tokens=15T, Train Compute (10^23 FLOPs)=7.22026.04 | 26.8 | 69.9 | 44.4 | 26.1 | |
| Qwen 3# Parameters=8B, # Train Tokens=36T, Train Compute (10^23 FLOPs)=17.32026.04 | 26.2 | 76.5 | 50.1 | 47.6 | |
| Olmo 3# Parameters=7B, # Train Tokens=6T, Train Compute (10^23 FLOPs)=2.6, Decoding Settings=--disable-cascade-attn, --enforce-eager2026.04 | 17.7 | 64 | 37.2 | 23.6 | |
| Olmo Hybrid# Parameters=7B, # Train Tokens=6T, Train Compute (10^23 FLOPs)=2.6, Decoding Settings=--disable-cascade-attn, --enforce-eager2026.04 | 16.8 | 65.2 | 41.7 | 23.4 | |
| Recurr. Gemma# Parameters=9B, # Train Tokens=2T, Train Compute (10^23 FLOPs)=1.02026.04 | 5.8 | 53.3 | 27.6 | 17.3 | |
| Falcon Mamba# Parameters=7B, # Train Tokens=6T, Train Compute (10^23 FLOPs)=2.52026.04 | 5.7 | 42.8 | 24.5 | 15.5 | |
| xLSTM# Parameters=7B, # Train Tokens=2T, Train Compute (10^23 FLOPs)=0.92026.04 | 0.6 | 24.6 | 11.4 | 7.5 |