Language Modeling on Lambada OpenAI
78.4AccuracyTWLA
Evaluation Results
| Method | Links | |
|---|---|---|
| TWLABackbone=LLaMA2-70B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 78.4 | |
| TWLABackbone=LLaMA2-13B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 76.4 | |
| BaselineModel=OLMoE, Weight Bits=16, Zero-shot=true2026.04 | 70.83 | |
| TWLABackbone=LLaMA2-7B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 70.33 | |
| Pythia (Deduplicated)Model size=12B, Evaluation protocol=Five-shot2023.04 | 69.1 | |
| TWLABackbone=Qwen3-14B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 69.05 | |
| Ponder-2.8BShots=0-shot, Training tokens=300B2026.03 | 68.9 | |
| BaselineModel=DeepSeekV2-Lite, Weight Bits=16, Zero-shot=true2026.04 | 68.37 | |
| AdaPonderLM-2.8BShots=0-shot, Training tokens=305B2026.03 | 68.3 | |
| PonderLM-2-Pythia-1.4BEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=1x2025.09 | 67.6 | |
| TWLABackbone=Qwen3-32B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 67.41 | |
| PythiaNumber of Parameters=12B, Evaluation Protocol=Five-shot2023.04 | 67.3 | |
| Pythia-6.9BShots=0-shot, Training tokens=300B2026.03 | 67.2 | |
| Pythia (Deduplicated)Model size=6.9B, Evaluation protocol=Five-shot2023.04 | 66.3 | |
| Ponder-1.4BEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=2x2025.09 | 65.2 | |
| Ponder-1.4BShots=0-shot, Training tokens=300B2026.03 | 65.2 | |
| TWLABackbone=LLaMA3-8B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 64.93 | |
| BaselineModel=Qwen3-MoE, Weight Bits=16, Zero-shot=true2026.04 | 64.86 | |
| Pythia-2.8BEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=1x2025.09 | 64.6 | |
| Pythia-2.8BShots=0-shot, Training tokens=300B2026.03 | 64.6 | |
| AdaPonderLM-1.4BShots=0-shot, Training tokens=312B2026.03 | 64.5 | |
| PonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 64.2 | |
| PythiaNumber of Parameters=6.9B, Evaluation Protocol=Five-shot2023.04 | 63.8 | |
| PonderLM-2-Pythia-1.4BEvaluation Protocol=5-shot, #training tokens=300B, Inference Cost=1x2025.09 | 63.6 | |
| OPT-2.7BShots=0-shot, Training tokens=300B2026.03 | 63.5 | |
| AdaPonderLM-2.8BShots=5-shot, Training tokens=300B2026.03 | 63.5 | |
| Pythia-6.9BShots=5-shot, Training tokens=300B2026.03 | 62.5 | |
| GPTneo-2.7BShots=0-shot, Training tokens=300B2026.03 | 62.1 | |
| Pythia-1.4BEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=1x2025.09 | 61.6 | |
| Pythia-1.4BShots=0-shot, Training tokens=300B2026.03 | 61.6 | |
| Pythia (Deduplicated)Model size=2.8B, Evaluation protocol=Five-shot2023.04 | 60.6 | |
| PythiaNumber of Parameters=2.8B, Evaluation Protocol=Five-shot2023.04 | 60.5 | |
| TWLABackbone=Qwen3-8B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 60.29 | |
| OPT-2.7BShots=5-shot, Training tokens=300B2026.03 | 60.2 | |
| PonderLM-1.4BShots=5-shot, Training tokens=300B2026.03 | 59.2 | |
| PonderLM-2-Pythia-410MEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=1x2025.09 | 59.1 | |
| AdaPonderLM-1.4BShots=5-shot, Training tokens=310B2026.03 | 59 | |
| Pythia-2.8BShots=5-shot, Training tokens=300B2026.03 | 59 | |
| Tinyllama-1.1BEvaluation Protocol=0-shot, #training tokens=3T, Inference Cost=1x2025.09 | 58.8 | |
| Tinyllama-1.1BShots=0-shot, Training tokens=3T2026.03 | 58.8 | |
| OPT-1.3BEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=1x2025.09 | 57.9 | |
| OPT-1.3BShots=0-shot, Training tokens=300B2026.03 | 57.9 | |
| PythiaNumber of Parameters=1.4B, Evaluation Protocol=Five-shot2023.04 | 57.8 | |
| Ponder-410MEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=2x2025.09 | 56.9 | |
| Pythia (Deduplicated)Model size=1.4B, Evaluation protocol=Five-shot2023.04 | 56.8 | |
| GPTneo-2.7BShots=5-shot, Training tokens=300B2026.03 | 56 | |
| Pythia-1BEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=1x2025.09 | 55.9 | |
| Pythia-1.4BShots=5-shot, Training tokens=300B2026.03 | 54.5 | |
| OPT-1.3BShots=5-shot, Training tokens=300B2026.03 | 54 | |
| Tinyllama-1.1BShots=5-shot, Training tokens=3T2026.03 | 53.8 | |
| ConceptLMModel Scale=410M, Backbone Architecture=Pythia, Zero-shot Evaluation=true2026.02 | 53 | |
| Pythia (Deduplicated)Model size=1B, Evaluation protocol=Five-shot2023.04 | 52.8 | |
| ContextLMModel Scale=410M, Backbone Architecture=Pythia, Zero-shot Evaluation=true2026.02 | 52.2 | |
| PonderLM-2-Pythia-410MEvaluation Protocol=5-shot, #training tokens=300B, Inference Cost=1x2025.09 | 52.1 | |
| MoBiEModel=Qwen3-MoE, Weight Bits=1.34, Zero-shot=true2026.04 | 52.01 | |
| Bloom-3BShots=0-shot, Training tokens=366B2026.03 | 51.7 | |
| Pythia-410MModel Scale=410M, Backbone Architecture=Pythia, Zero-shot Evaluation=true2026.02 | 51.6 | |
| Pythia-410MEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=1x2025.09 | 51.4 | |
| Pythia-410MShots=0-shot, Training tokens=300B2026.03 | 51.4 | |
| PythiaNumber of Parameters=1B, Evaluation Protocol=Five-shot2023.04 | 50.7 | |
| GPTQModel=Qwen3-MoE, Weight Bits=3, Zero-shot=true2026.04 | 48.22 | |
| MoBiEModel=OLMoE, Weight Bits=1.46, Zero-shot=true2026.04 | 47.12 | |
| Pythia (Deduplicated)Model size=410M, Evaluation protocol=Five-shot2023.04 | 46.6 | |
| Pause Token-410MShots=0-shot, Training tokens=25B2026.03 | 46.3 | |
| Bloom-1.7BEvaluation Protocol=0-shot, #training tokens=366B, Inference Cost=1x2025.09 | 46.2 | |
| Bloom-1.7BShots=0-shot, Training tokens=366B2026.03 | 46.2 | |
| Bloom-3BShots=5-shot, Training tokens=366B2026.03 | 46.2 | |
| PonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 45.8 | |
| PythiaNumber of Parameters=410M, Evaluation Protocol=Five-shot2023.04 | 45.5 | |
| AdaPonderLM-410MShots=0-shot, Training tokens=25B2026.03 | 45.4 | |
| OPT-350MEvaluation Protocol=0-shot, #training tokens=300B, Inference Cost=1x2025.09 | 45.2 | |
| GPTQModel=OLMoE, Weight Bits=3, Zero-shot=true2026.04 | 45.12 | |
| Loop Transformer-410MShots=0-shot, Training tokens=25B2026.03 | 44.3 | |
| Pythia-410MEvaluation Protocol=5-shot, #training tokens=300B, Inference Cost=1x2025.09 | 43.9 | |
| Pythia-410MShots=5-shot, Training tokens=300B2026.03 | 43.9 | |
| NoWagModel=Qwen3-MoE, Weight Bits=2.11, Zero-shot=true2026.04 | 42.99 | |
| ConceptLMModel Scale=1.5B, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 42.9 | |
| Bloom-1.7BShots=5-shot, Training tokens=366B2026.03 | 42.5 | |
| ContextLMModel Scale=1.5B, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 41.9 | |
| ConceptLMModel Scale=160M, Backbone Architecture=Pythia, Zero-shot Evaluation=true2026.02 | 41.8 | |
| ConceptLMModel Scale=774M, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 41.1 | |
| ContextLMModel Scale=774M, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 40.7 | |
| Pause Token-410MShots=5-shot, Training tokens=25B2026.03 | 40.3 | |
| PonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 39.7 | |
| Loop Transformer-410MShots=5-shot, Training tokens=25B2026.03 | 39.6 | |
| GPT-1.5BModel Scale=1.5B, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 39.3 | |
| GPT-PMModel Scale=1.5B, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 38.7 | |
| AdaPonderLM-410MShots=5-shot, Training tokens=25B2026.03 | 38.7 | |
| GPT-774MModel Scale=774M, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 37.2 | |
| GPT-PMModel Scale=774M, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 37.2 | |
| ContextLMModel Scale=160M, Backbone Architecture=Pythia, Zero-shot Evaluation=true2026.02 | 37.2 | |
| ContextLMModel Scale=355M, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 36.8 | |
| ConceptLMModel Scale=355M, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 36.6 | |
| MoBiEModel=DeepSeekV2-Lite, Weight Bits=1.47, Zero-shot=true2026.04 | 35.98 | |
| GPT-PMModel Scale=355M, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 35.2 | |
| Bloom-560MEvaluation Protocol=0-shot, #training tokens=366B, Inference Cost=1x2025.09 | 34.3 | |
| GPTQModel=DeepSeekV2-Lite, Weight Bits=3, Zero-shot=true2026.04 | 33.22 | |
| AdamWModel Scale=1.2B, Pre-training Budget=1x Chinchilla, Evaluation Framework=LM Eval-Harness2024.11 | 33.11 | |
| GPT-355MModel Scale=355M, Backbone Architecture=GPT-2, Zero-shot Evaluation=true2026.02 | 32.8 | |
| Pythia-160MModel Scale=160M, Backbone Architecture=Pythia, Zero-shot Evaluation=true2026.02 | 32.7 |