Mean Performance Evaluation on Downstream Tasks Summary
61.5Average AccuracyPonderLM-2.8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 61.5 | 3.9 | |
| AdaPonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 61.1 | 3.5 | |
| Ponder-2.8BShots=0-shot, Training tokens=300B2026.03 | 60.4 | 3.1 | |
| Pythia-6.9BShots=5-shot, Training tokens=300B2026.03 | 60.1 | — | |
| AdaPonderLM-2.8BShots=0-shot, Training tokens=305B2026.03 | 59.6 | 2.2 | |
| Pythia-6.9BShots=0-shot, Training tokens=300B2026.03 | 59.1 | — | |
| OPT-2.7BShots=5-shot, Training tokens=300B2026.03 | 58.2 | — | |
| PonderLM-1.4BShots=5-shot, Training tokens=300B2026.03 | 57.6 | 3.6 | |
| Pythia-2.8BShots=5-shot, Training tokens=300B2026.03 | 57.6 | — | |
| Pythia-2.8BShots=0-shot, Training tokens=300B2026.03 | 57.3 | — | |
| AdaPonderLM-1.4BShots=5-shot, Training tokens=310B2026.03 | 56.8 | 2.8 | |
| OPT-2.7BShots=0-shot, Training tokens=300B2026.03 | 56.7 | — | |
| Ponder-1.4BShots=0-shot, Training tokens=300B2026.03 | 56.5 | 2.4 | |
| GPTneo-2.7BShots=5-shot, Training tokens=300B2026.03 | 56.3 | — | |
| AdaPonderLM-1.4BShots=0-shot, Training tokens=312B2026.03 | 55.9 | 1.8 | |
| Tinyllama-1.1BShots=5-shot, Training tokens=3T2026.03 | 55.9 | — | |
| GPTneo-2.7BShots=0-shot, Training tokens=300B2026.03 | 55.5 | — | |
| Tinyllama-1.1BShots=0-shot, Training tokens=3T2026.03 | 55.4 | — | |
| Pythia-1.4BShots=0-shot, Training tokens=300B2026.03 | 54.1 | — | |
| Pythia-1.4BShots=5-shot, Training tokens=300B2026.03 | 54.1 | — | |
| Bloom-3BShots=5-shot, Training tokens=366B2026.03 | 54.1 | — | |
| Bloom-3BShots=0-shot, Training tokens=366B2026.03 | 53.9 | — | |
| OPT-1.3BShots=0-shot, Training tokens=300B2026.03 | 53.6 | — | |
| OPT-1.3BShots=5-shot, Training tokens=300B2026.03 | 52.7 | — | |
| Bloom-1.7BShots=5-shot, Training tokens=366B2026.03 | 50.9 | — | |
| Bloom-1.7BShots=0-shot, Training tokens=366B2026.03 | 50.2 | — | |
| Pythia-410MShots=0-shot, Training tokens=300B2026.03 | 47.6 | — | |
| Pythia-410MShots=5-shot, Training tokens=300B2026.03 | 47.6 | — | |
| Pause Token-410MShots=0-shot, Training tokens=25B2026.03 | 46.1 | — | |
| PonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 45.9 | — | |
| Pause Token-410MShots=5-shot, Training tokens=25B2026.03 | 45.7 | — | |
| PonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 45.3 | — | |
| Loop Transformer-410MShots=5-shot, Training tokens=25B2026.03 | 45.1 | — | |
| AdaPonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 44.8 | — | |
| Loop Transformer-410MShots=0-shot, Training tokens=25B2026.03 | 44.7 | — | |
| AdaPonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 44.3 | — |