Commonsense Reasoning on Winogrande (Score)
77.74ScoreGemma3
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemma3Architecture=Dense, # Total Params=12.0B, # Trained Tokens=12T2025.10 | 77.74 | |
| Llama3.1Architecture=Dense, # Total Params=8.0B, # Trained Tokens=15T2025.10 | 77.11 | |
| Qwen3Architecture=Dense, # Total Params=8.0B, # Trained Tokens=36T2025.10 | 76.8 | |
| Qwen2.5Architecture=Dense, # Total Params=7.0B, # Trained Tokens=18T2025.10 | 76.48 | |
| Ouro 2.6B R4Architecture=LoopLM, # Total Params=2.6B, # Trained Tokens=7.7T2025.10 | 75.85 | |
| Gemma3-4BModel Scale=4B2026.04 | 72.77 | |
| Llama3.2-3BModel Scale=3B2026.04 | 72.38 | |
| Gemma3Architecture=Dense, # Total Params=4.0B, # Trained Tokens=4T2025.10 | 71.27 | |
| Qwen3Architecture=Dense, # Total Params=4.0B, # Trained Tokens=36T2025.10 | 71.19 | |
| Qwen3-4BModel Scale=4B2026.04 | 70.48 | |
| Qwen2.5Architecture=Dense, # Total Params=3.0B, # Trained Tokens=18T2025.10 | 70.17 | |
| Qwen2.5-3BModel Scale=3B2026.04 | 69.46 | |
| Llama3.2Architecture=Dense, # Total Params=3.0B, # Trained Tokens=9T2025.10 | 69.14 | |
| SpB2.0-5BModel Scale=5B, Additional Training Tokens=14B2026.04 | 68.35 | |
| Pre-LNArchitecture=512 Layers / 3B Params, Peak Learning Rate=3.0 x 10^-3, Pre-training Tokens=1T, Evaluation Protocol (Shots)=0-shot2026.01 | 66.7 | |
| KEELArchitecture=512 Layers / 3B Params, Peak Learning Rate=4.5 x 10^-3, Pre-training Tokens=1T, Evaluation Protocol (Shots)=0-shot2026.01 | 66.7 | |
| DPModel Scale=Gemma 9B, Method=DP2026.04 | 64.2 | |
| Decoupled DiLoCoModel Scale=Gemma 9B, Method=DiLoCo2026.04 | 63.8 | |
| Decoupled DiLoCoModel Scale=Gemma 5B, Method=DiLoCo2026.04 | 61.3 | |
| DPModel Scale=Gemma 5B, Method=DP2026.04 | 59.6 | |
| DPModel Scale=Gemma 2B, Method=DP2026.04 | 55.6 | |
| Decoupled DiLoCoModel Scale=Gemma 2B, Method=DiLoCo2026.04 | 54.1 |