General Language Understanding on General Downstream Tasks Aggregate
59.5Average AccuracyPonderLM-2-Pythia-1.4B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PonderLM-2-Pythia-1.4BEvaluation Protocol=5-shot, #training tokens=300B2025.09 | 59.5 | 5.4 | |
| PonderLM-2-Pythia-1.4BEvaluation Protocol=0-shot, #training tokens=300B2025.09 | 58.5 | 4.4 | |
| AuroraModel Size=1.1B2026.06 | 57.2 | — | |
| Ponder-1.4BEvaluation Protocol=0-shot, #training tokens=300B2025.09 | 56.5 | — | |
| MuonModel Size=1.1B2026.06 | 56 | — | |
| U-NorMuonModel Size=1.1B2026.06 | 55.5 | — | |
| NorMuonModel Size=1.1B2026.06 | 55.1 | — | |
| Pythia-1.4BEvaluation Protocol=0-shot, #training tokens=300B2025.09 | 54.1 | — | |
| PonderLM-2-Pythia-410MEvaluation Protocol=0-shot, #training tokens=300B2025.09 | 51.9 | 4.3 | |
| PonderLM-2-Pythia-410MEvaluation Protocol=5-shot, #training tokens=300B2025.09 | 51.9 | 4.3 | |
| Ponder-410MEvaluation Protocol=0-shot, #training tokens=300B2025.09 | 50.4 | — | |
| Pythia-410MEvaluation Protocol=0-shot, #training tokens=300B2025.09 | 47.6 | — | |
| AuroraModel Size=340M2026.06 | 46.4 | — | |
| U-NorMuonModel Size=340M2026.06 | 45.8 | — | |
| NorMuonModel Size=340M2026.06 | 45.7 | — | |
| MuonModel Size=340M2026.06 | 45.6 | — |