Word Prediction on LAMBADA (test)
87.15AccuracyMT-NLG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MT-NLGshot_setting=few-shot2022.01 | 87.15 | — | |
| GPT-3Parameters (M)=1750002020.04 | 86.4 | — | |
| GPT-3shot_setting=few-shot2022.01 | 86.4 | — | |
| Humanestimation_source=100 randomly-sampled DEV instances2016.10 | 86 | — | |
| Human performance2020.04 | 86 | — | |
| MT-NLGshot_setting=zero-shot2022.01 | 76.56 | — | |
| GPT-3shot_setting=zero-shot2022.01 | 76.2 | — | |
| Gophershot_setting=zero-shot2022.01 | 74.5 | — | |
| MT-NLGshot_setting=one-shot2022.01 | 73.06 | — | |
| GPT-3shot_setting=one-shot2022.01 | 72.5 | — | |
| Universal TransformerParameters (M)=1522020.04 | 56 | — | |
| RSEParameters (M)=112020.04 | 54.34 | — | |
| SEParameters (M)=332020.04 | 52.28 | — | |
| Mistral (Full-Attention)Model Scale=1.4B, Evaluation Protocol=Zero-Shot2024.07 | 50.1 | — | |
| GA Reader + featuresfeatures=Wang et al. (2016)2016.10 | 49 | — | |
| Gated-Attention ReaderParameters (M)=unknown2020.04 | 49 | — | |
| PGM 8 / 8distillation=After Distillation (SDTT)2025.05 | 47.22 | — | |
| PGM 8 / 8distillation=Before2025.05 | 46.98 | — | |
| GA Reader2016.10 | 45.4 | — | |
| BMoJo (Fading)Model Scale=1.4B, Evaluation Protocol=Zero-Shot2024.07 | 45.4 | — | |
| BMoJo (Fading + Eidetic)Model Scale=1.4B, Evaluation Protocol=Zero-Shot2024.07 | 44.8 | — | |
| AS Reader + featuresfeatures=Wang et al. (2016)2016.10 | 44.5 | — | |
| PGM 6 / 6 (1024)distillation=After Distillation (SDTT)2025.05 | 44.48 | — | |
| Mamba (SSM)Model Scale=1.4B, Evaluation Protocol=Zero-Shot2024.07 | 43.9 | — | |
| AS Reader2016.10 | 41.4 | — | |
| PGM 6 / 6 (1024)distillation=Before2025.05 | 41.39 | — | |
| MDLMdistillation=After Distillation (SDTT)2025.05 | 41.34 | — | |
| MDLMdistillation=Before2025.05 | 38.52 | — | |
| Hybrid (Sliding Attention + SSM)Model Scale=1.4B, Evaluation Protocol=Zero-Shot2024.07 | 37.6 | — | |
| Modified Stanford Reader2016.10 | 32.1 | — | |
| Mistral (Full-Attention)Model Scale=370M, Evaluation Protocol=Zero-Shot2024.07 | 31.6 | — | |
| Mamba (SSM)Model Scale=370M, Evaluation Protocol=Zero-Shot2024.07 | 31.4 | — | |
| BMoJo (Fading)Model Scale=370M, Evaluation Protocol=Zero-Shot2024.07 | 29.6 | — | |
| BMoJo (Fading + Eidetic)Model Scale=370M, Evaluation Protocol=Zero-Shot2024.07 | 28.6 | — | |
| Hybrid (Sliding Attention + SSM)Model Scale=370M, Evaluation Protocol=Zero-Shot2024.07 | 26.3 | — | |
| Stanford Reader2016.10 | 21.7 | — | |
| GPT-2Parameters=124M, zero-shot=true2024.01 | 17.1 | — | |
| PIXARParameters=113M, Stage=2, zero-shot=true2024.01 | 13.8 | 82.2 | |
| n-gram + cachebaseline_category=Context-restricted language model2016.10 | 11.8 | — | |
| Most frequentbaseline_category=Context-restricted non-stopword2016.10 | 11.7 | — | |
| n-grambaseline_category=Context-restricted language model2016.10 | 10.7 | — | |
| LSTMbaseline_category=Context-restricted language model2016.10 | 9.2 | — | |
| Random cap. in contextbaseline_category=Paperno et al. (2016)2016.10 | 7.3 | — | |
| Lastbaseline_category=Context-restricted non-stopword2016.10 | 6.2 | — | |
| PIXARParameters=113M, Stage=1, zero-shot=true2024.01 | 5.7 | 54.8 | |
| Randombaseline_category=Context-restricted non-stopword2016.10 | 5.6 | — | |
| Firstbaseline_category=Context-restricted non-stopword2016.10 | 3.8 | — | |
| Random in contextbaseline_category=Paperno et al. (2016)2016.10 | 1.6 | — | |
| Random word2020.04 | 1.6 | — | |
| n-grambaseline_category=Paperno et al. (2016)2016.10 | 0.1 | — | |
| n-gram + cachebaseline_category=Paperno et al. (2016)2016.10 | 0.1 | — | |
| LSTMbaseline_category=Paperno et al. (2016)2016.10 | 0 | — | |
| Memory networkbaseline_category=Paperno et al. (2016)2016.10 | 0 | — |