Language Modeling Accuracy on Lambada Standard
60.8AccuracyPonder-2.8B
Evaluation Results
| Method | Links | |
|---|---|---|
| Ponder-2.8BShots=0-shot, Training tokens=300B2026.03 | 60.8 | |
| AdaPonderLM-2.8BShots=0-shot, Training tokens=305B2026.03 | 59.8 | |
| PonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 58.7 | |
| AdaPonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 58 | |
| OPT-2.7BShots=0-shot, Training tokens=300B2026.03 | 56 | |
| Pythia-6.9BShots=0-shot, Training tokens=300B2026.03 | 55.9 | |
| OPT-2.7BShots=5-shot, Training tokens=300B2026.03 | 55 | |
| Pythia-6.9BShots=5-shot, Training tokens=300B2026.03 | 54.8 | |
| Pythia-2.8BShots=0-shot, Training tokens=300B2026.03 | 54.3 | |
| Ponder-1.4BShots=0-shot, Training tokens=300B2026.03 | 53.8 | |
| OPT-1.3BShots=0-shot, Training tokens=300B2026.03 | 52.5 | |
| AdaPonderLM-1.4BShots=0-shot, Training tokens=312B2026.03 | 52.4 | |
| GPTneo-2.7BShots=0-shot, Training tokens=300B2026.03 | 51.6 | |
| GPTneo-2.7BShots=5-shot, Training tokens=300B2026.03 | 51.6 | |
| Bloom-3BShots=0-shot, Training tokens=366B2026.03 | 50.9 | |
| Pythia-2.8BShots=5-shot, Training tokens=300B2026.03 | 50.7 | |
| PonderLM-1.4BShots=5-shot, Training tokens=300B2026.03 | 49.9 | |
| Pythia-1.4BShots=0-shot, Training tokens=300B2026.03 | 49.7 | |
| Tinyllama-1.1BShots=0-shot, Training tokens=3T2026.03 | 49.3 | |
| OPT-1.3BShots=5-shot, Training tokens=300B2026.03 | 49 | |
| AdaPonderLM-1.4BShots=5-shot, Training tokens=310B2026.03 | 48.1 | |
| Bloom-3BShots=5-shot, Training tokens=366B2026.03 | 47.1 | |
| Tinyllama-1.1BShots=5-shot, Training tokens=3T2026.03 | 45 | |
| Bloom-1.7BShots=0-shot, Training tokens=366B2026.03 | 44.5 | |
| Pythia-1.4BShots=5-shot, Training tokens=300B2026.03 | 44.5 | |
| Bloom-1.7BShots=5-shot, Training tokens=366B2026.03 | 41.5 | |
| Pythia-410MShots=0-shot, Training tokens=300B2026.03 | 36.4 | |
| Pause Token-410MShots=0-shot, Training tokens=25B2026.03 | 35.7 | |
| PonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 33.6 | |
| PonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 33.5 | |
| Pythia-410MShots=5-shot, Training tokens=300B2026.03 | 32.8 | |
| Pause Token-410MShots=5-shot, Training tokens=25B2026.03 | 32.8 | |
| AdaPonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 32.6 | |
| Loop Transformer-410MShots=0-shot, Training tokens=25B2026.03 | 31.9 | |
| Loop Transformer-410MShots=5-shot, Training tokens=25B2026.03 | 29.7 | |
| AdaPonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 29.5 |