Language Modeling on GovReport
2.5PerplexityPoSE
Evaluation Results
| Method | Links | |
|---|---|---|
| PoSEContext Length=32K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 2.5 | |
| CREAMContext Length=32K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 2.5 | |
| PoSEContext Length=32K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 2.5 | |
| CREAMContext Length=32K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 2.5 | |
| PoSEContext Length=32K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 2.6 | |
| PoSEContext Length=16K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 2.7 | |
| CREAMContext Length=16K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 2.7 | |
| PoSEContext Length=16K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 2.7 | |
| CREAMContext Length=16K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 2.7 | |
| CREAMContext Length=32K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 2.7 | |
| PoSEContext Length=16K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 2.7 | |
| CREAMContext Length=16K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 2.7 | |
| PoSEContext Length=8K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 3.2 | |
| CREAMContext Length=8K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 3.2 | |
| PoSEContext Length=8K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 3.2 | |
| CREAMContext Length=8K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 3.2 | |
| PoSEContext Length=8K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 3.2 | |
| CREAMContext Length=8K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 3.2 | |
| LLaMa2-7BContext Length=4K, Fine-tuned Length=4K2024.06 | 3.6 | |
| RandPosContext Length=16K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 3.6 | |
| PoSEContext Length=4K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 3.7 | |
| PoSEContext Length=4K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 3.7 | |
| CREAMContext Length=4K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 3.7 | |
| PoSEContext Length=4K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 3.8 | |
| CREAMContext Length=4K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 3.8 | |
| CREAMContext Length=4K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 3.8 | |
| RandPosContext Length=8K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 4 | |
| RandPosContext Length=32K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 4 | |
| RandPosContext Length=16K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 4 | |
| RandPosContext Length=8K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 4.4 | |
| RandPosContext Length=4K, Interpolation Method=NTK, Fine-tuned Length=4K2024.06 | 4.6 | |
| RandPosContext Length=32K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 4.6 | |
| RandPosContext Length=4K, Interpolation Method=YaRN, Fine-tuned Length=4K2024.06 | 5 | |
| RandPosContext Length=32K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 5.8 | |
| RandPosContext Length=16K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 6.2 | |
| RandPosContext Length=8K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 7.4 | |
| RandPosContext Length=4K, Interpolation Method=Linear, Fine-tuned Length=4K2024.06 | 8.9 | |
| (w+a)kNN-LMa (rescore)LMa=domain-adapted GPT-2, datastore=combined, rescore=true, context_representation=LN22022.11 | 12.79 | |
| (a)kNN-LMa (rescore)LMa=domain-adapted GPT-2, datastore=adaptation, rescore=true, context_representation=LN22022.11 | 12.87 | |
| (w+a)kNN-LMaLMa=domain-adapted GPT-2, datastore=combined, context_representation=LN22022.11 | 13.01 | |
| (a)kNN-LMaLMa=domain-adapted GPT-2, datastore=adaptation, context_representation=LN22022.11 | 13.08 | |
| (w)kNN-LMa (rescore)LMa=domain-adapted GPT-2, datastore=pretraining, rescore=true, context_representation=LN22022.11 | 14.36 | |
| (w)kNN-LMaLMa=domain-adapted GPT-2, datastore=pretraining, context_representation=LN22022.11 | 14.47 | |
| LMa (only)LMa=domain-adapted GPT-22022.11 | 14.72 | |
| (a)kNN-LMLM=GPT-2, datastore=adaptation, context_representation=original2022.11 | 14.87 | |
| (w)kNN-LMLM=GPT-2, datastore=pretraining, context_representation=LN22022.11 | 18.81 | |
| (w)kNN-LMLM=GPT-2, datastore=pretraining, context_representation=original2022.11 | 18.99 | |
| LM (only)LM=GPT-22022.11 | 19.32 | |
| (a)kNN (only)datastore=adaptation2022.11 | 40.39 | |
| (w)kNN (only)datastore=pretraining2022.11 | 83.62 |