Word Prediction Accuracy on LAMBADA
86.9AccuracyPaLM 2-L
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PaLM 2-Lprompting=1-shot2023.05 | 86.9 | — | |
| GPT-3setting=few-shot, model_size=175B2020.05 | 86.4 | 1.92 | |
| PaLM 2-Mprompting=1-shot2023.05 | 83.7 | — | |
| PaLMprompting=1-shot2023.05 | 81.8 | — | |
| PaLM 2-Sprompting=1-shot2023.05 | 80.7 | — | |
| GPT-3setting=zero-shot, model_size=175B2020.05 | 76.2 | 3 | |
| Our Trained ModelModel Size=52B, Data type=W8A8, Zero-shot=true2023.05 | 75.5 | — | |
| EnsembleAggregation Method=Byte-level Ensemble2025.06 | 75.5 | — | |
| Our Trained ModelModel Size=52B, Data type=FP16, Zero-shot=true2023.05 | 75.47 | — | |
| GPT-3Model Variant=DaVinci, Zero-shot=true2022.04 | 75.2 | — | |
| FP16Rank=-, Zero-shot=true, Model=Vicuna-13B2026.03 | 74.33 | — | |
| ZeroQuant-V2Zero-shot=true, Model=Vicuna-13B2026.03 | 74 | — | |
| L2QERRank=64, Zero-shot=true, Model=Vicuna-13B2026.03 | 74 | — | |
| GlowQZero-shot=true, Model=Vicuna-13B2026.03 | 74 | — | |
| Llama2 7BCR=None2025.09 | 73.7 | — | |
| PKU-RLHF aligned model (M_PKU-RLHF)Model Family=gemma-2-2b-it2026.03 | 73.04 | — | |
| QERAZero-shot=true, Model=Vicuna-13B2026.03 | 73 | — | |
| QWEN3Model=QWEN32025.06 | 72.7 | — | |
| GPT-3setting=one-shot, model_size=175B2020.05 | 72.5 | 3.35 | |
| GPT-NeoXModel Size=20B, Zero-shot=true2022.04 | 72 | — | |
| CoSpaDi (grouped)CR=0.22025.09 | 71.3 | — | |
| FairSeqNumber of Parameters=13B, Zero-Shot=true2022.04 | 70.9 | — | |
| Our Trained ModelModel Size=13B, Data type=FP16, Zero-shot=true2023.05 | 70.81 | — | |
| Our Trained ModelModel Size=6B, Data type=FP16, Zero-shot=true2023.05 | 70.5 | — | |
| CoSpaDi (per-layer)CR=0.22025.09 | 70.2 | — | |
| Our Trained ModelModel Size=6B, Data type=W8A8, Zero-shot=true2023.05 | 70 | — | |
| Our Trained ModelModel Size=13B, Data type=W8A8, Zero-shot=true2023.05 | 69.9 | — | |
| GPT-NeoXParameters=20B, Evaluation=Five-shot2022.04 | 69.8 | — | |
| GPT-3Model Variant=Curie, Zero-shot=true2022.04 | 69.3 | — | |
| DensePrune Rate=0%, Evaluation Protocol=zero-shot2024.06 | 69.14 | — | |
| NHA-Qwen3-30BA3BModel=NHA-Qwen3-30BA3B, Architecture Type=Native Hybrid Attention, Full-attention layers=32025.10 | 68.85 | — | |
| LoopQModel=Ouro 2.6B, Quantization=W4A82026.05 | 68.37 | 4.47 | |
| GPT-JModel Size=6B, Zero-shot=true2022.04 | 68.3 | — | |
| SOTA2020.05 | 68 | 8.63 | |
| Safety-Reset model (M_base)Model Family=gemma-2-2b-it2026.03 | 67.61 | — | |
| BaseModel=LLaMA-3.1-8B, Compression=0%, Mode=Zero-shot2025.10 | 67.36 | — | |
| FairSeqNumber of Parameters=6.7B, Zero-Shot=true2022.04 | 67.3 | — | |
| Our Trained ModelModel Size=52B, Data type=W4, Zero-shot=true2023.05 | 66.85 | — | |
| GPT-JParameters=6B, Evaluation=Five-shot2022.04 | 66.2 | — | |
| Self-MOA aligned model (M_Self-MOA)Model Family=gemma-2-2b-it2026.03 | 65.46 | — | |
| GlowQ-SZero-shot=true, Model=Vicuna-13B2026.03 | 65.33 | — | |
| BF16Model=Ouro 1.4B, Quantization=W4A42026.05 | 65.05 | 5.3 | |
| CoSpaDi (grouped)CR=0.32025.09 | 64.5 | — | |
| Original model (M_original)Model Family=gemma-2-2b-it2026.03 | 64.04 | — | |
| SVD-LLMCR=0.22025.09 | 64 | — | |
| FairSeqNumber of Parameters=2.7B, Zero-Shot=true2022.04 | 63.2 | — | |
| Qwen3-30BA3BModel=Qwen3-30BA3B, Architecture Type=Transformer2025.10 | 63.05 | — | |
| Basis SharingCR=0.22025.09 | 63 | — | |
| Llama-3.2Size=1.2B2025.04 | 62.99 | — | |
| KALAVAI MoEModel Scale=Pythia-6.9B, Training Steps=10000 base, Freeze Layers=6, Seed=42, Number of Examples=5002026.03 | 62.8 | — | |
| OLMO2Model=OLMO22025.06 | 62.8 | — | |
| CoSpaDi (per-layer)CR=0.32025.09 | 62.7 | — | |
| PG pruningPrune Rate=10%, Evaluation Protocol=zero-shot2024.06 | 62.63 | — | |
| LoopQModel=Ouro 1.4B, Quantization=W4A82026.05 | 62.55 | 6.09 | |
| GPT-3Model Variant=Babbage, Zero-shot=true2022.04 | 62.5 | — | |
| Depth-AttentionSize=3B, Evaluation Protocol=Zero-shot2026.06 | 62.25 | — | |
| AverageAggregation Method=Average2025.06 | 62.2 | — | |
| mHCSize=3B, Evaluation Protocol=Zero-shot2026.06 | 61.73 | — | |
| Attention ResidualsSize=3B, Evaluation Protocol=Zero-shot2026.06 | 61.6 | — | |
| Base modelModel Scale=Pythia-6.9B, Training Steps=10000 base, Freeze Layers=6, Seed=42, Number of Examples=5002026.03 | 61.2 | — | |
| Original model (M_original)Model Family=qwen-2.5-1.5b-it2026.03 | 61.03 | — | |
| Safety-Reset model (M_base)Model Family=llama-3.2-1b-it2026.03 | 60.41 | — | |
| Original model (M_original)Model Family=llama-3.2-1b-it2026.03 | 60.24 | — | |
| Base modelBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 60.2 | — | |
| PKU-RLHF aligned model (M_PKU-RLHF)Model Family=llama-3.2-1b-it2026.03 | 60.16 | — | |
| Safety-Reset model (M_base)Model Family=qwen-2.5-1.5b-it2026.03 | 60.08 | — | |
| DenseFormerSize=3B, Evaluation Protocol=Zero-shot2026.06 | 59.91 | — | |
| Depth-AttentionSize=1.5B, Evaluation Protocol=Zero-shot2026.06 | 59.89 | — | |
| GPT-3 (SLW 8x Bsz)Model size=1.3B, few-shot (k)=152021.08 | 59.7 | — | |
| SliceGPTPrune Rate=10%, Evaluation Protocol=zero-shot2024.06 | 59.67 | — | |
| Self-MOA aligned model (M_Self-MOA)Model Family=qwen-2.5-1.5b-it2026.03 | 59.32 | — | |
| AMD-OLMoSize=1.2B2025.04 | 59.31 | — | |
| CLIMBSize=950M2025.04 | 59.05 | — | |
| KALAVAI MoEBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 59 | — | |
| VanillaSize=3B, Evaluation Protocol=Zero-shot2026.06 | 58.86 | — | |
| TinyLlamaSize=1.1B2025.04 | 58.84 | — | |
| GPT-3 (Baseline repro)Model size=1.3B, few-shot (k)=152021.08 | 58.8 | — | |
| PKU-RLHF aligned model (M_PKU-RLHF)Model Family=qwen-2.5-1.5b-it2026.03 | 58.8 | — | |
| Fict. specialistBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 58.6 | — | |
| Attention ResidualsSize=1.5B, Evaluation Protocol=Zero-shot2026.06 | 58.49 | — | |
| MonolithicBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 58.2 | — | |
| Weight averageBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 57.8 | — | |
| mHCSize=1.5B, Evaluation Protocol=Zero-shot2026.06 | 57.77 | — | |
| Code specialistBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 57.4 | — | |
| Depth-AttentionSize=3B, Evaluation Protocol=Five-shot2026.06 | 57.4 | — | |
| GPT-3 (Original)Model size=1.3B, few-shot (k)=152021.08 | 57 | — | |
| Self-MOA aligned model (M_Self-MOA)Model Family=llama-3.2-1b-it2026.03 | 56.88 | — | |
| LoopQModel=Ouro 2.6B, Quantization=W4A42026.05 | 56.8 | 8.17 | |
| mHCSize=3B, Evaluation Protocol=Five-shot2026.06 | 56.76 | — | |
| LoopQModel=Ouro 1.4B, Quantization=W4A42026.05 | 56.59 | 8.18 | |
| Sci. specialistBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 56.4 | — | |
| DenseFormerSize=1.5B, Evaluation Protocol=Zero-shot2026.06 | 56.34 | — | |
| FairSeqNumber of Parameters=1.3B, Zero-Shot=true2022.04 | 56.2 | — | |
| Attention ResidualsSize=3B, Evaluation Protocol=Five-shot2026.06 | 55.73 | — | |
| Our Trained ModelModel Size=6B, Data type=W4, Zero-shot=true2023.05 | 55.4 | — | |
| VanillaSize=1.5B, Evaluation Protocol=Zero-shot2026.06 | 55.02 | — | |
| BonsaiPrune Rate=10%, Evaluation Protocol=zero-shot2024.06 | 54.12 | — | |
| DenseFormerSize=3B, Evaluation Protocol=Five-shot2026.06 | 54.12 | — | |
| Depth-AttentionSize=1.5B, Evaluation Protocol=Five-shot2026.06 | 54.05 | — | |
| LLM-PrunerPrune Rate=10%, Evaluation Protocol=zero-shot2024.06 | 53.85 | — |